Repository navigation
[AWS OTel] increase cloudwatch autodiscover limit and add section to readme explaining the consequence - #21199
Conversation
…aining the consequence
✅ Elastic Docs Style Checker (Vale)No issues found on modified lines! The Vale linter checks documentation changes against the Elastic Docs style guide. To use Vale locally or report issues, refer to Elastic style guide for Vale. |
There was a problem hiding this comment.
🟡 Changes recommended
Documentation issues and the placeholder changelog URL must be corrected before approval.
Get a fresh assessment by requesting another Copilot review.
Pull request overview
This PR increases AWS CloudWatch OTel autodiscovery limits and documents the related metric and cost implications.
Changes:
- Raises selected autodiscovery defaults to 10,000.
- Adds limit and cost guidance to both README copies.
- Updates the package version and changelog.
File summaries
| File | Summary |
|---|---|
packages/aws_cloudwatch_input_otel/manifest.yml |
Updates package version and autodiscovery limits. |
packages/aws_cloudwatch_input_otel/docs/README.md |
Adds guidance; heading formatting, limit wording, and cost estimate require correction. |
packages/aws_cloudwatch_input_otel/changelog.yml |
Adds the enhancement entry; the placeholder PR URL requires replacement. |
packages/aws_cloudwatch_input_otel/_dev/build/docs/README.md |
Mirrors README updates; the same formatting, limit wording, and cost estimate issues require correction. |
Review details
Suppressed comments (6)
packages/aws_cloudwatch_input_otel/_dev/build/docs/README.md:88
- The receiver is capped by
discovery.limit, so a namespace with more metrics than the configured limit will not have all of its metrics collected. Please qualify this sentence to say metrics are collected up to the limit in the source and regenerate this copy.
This integration automatically discovers and collects all CloudWatch metrics published for the configured namespaces. The CloudWatch API bills per metric requested, so the more resources that exist in your account, the more metrics are discovered and the higher the collection cost.
packages/aws_cloudwatch_input_otel/_dev/build/docs/README.md:90
- This says the limit defaults to 10,000 for the integration, but the ALB, CLB, NLB, and GWLB policy templates still default to 250 (
manifest.yml:344,404,466,526). Please document the per-stream defaults or update the remaining templates so this cost guidance matches the actual configuration in the source and regenerate this copy.
To avoid unexpectedly large bills, the number of metrics collected is capped by the **Autodiscover Limit**, which defaults to 10,000, high enough to capture all metrics in a typical deployment. The limit applies to each namespace separately. It is a ceiling, not a fixed value - you are billed only for the metrics that actually exist, so smaller accounts cost proportionally less. With the 10,000 limit, the maximum cost per namespace is ~$3,500 per month at a 5-minute collection interval.
packages/aws_cloudwatch_input_otel/_dev/build/docs/README.md:90
- The generated README repeats the cost estimate that does not match the documented billing model and the receiver's three default statistics. Correct or remove the estimate here when updating the source README.
To avoid unexpectedly large bills, the number of metrics collected is capped by the **Autodiscover Limit**, which defaults to 10,000, high enough to capture all metrics in a typical deployment. The limit applies to each namespace separately. It is a ceiling, not a fixed value - you are billed only for the metrics that actually exist, so smaller accounts cost proportionally less. With the 10,000 limit, the maximum cost per namespace is ~$3,500 per month at a 5-minute collection interval.
packages/aws_cloudwatch_input_otel/docs/README.md:89
- The receiver is capped by
discovery.limit, so a namespace with more metrics than the configured limit will not have all of its metrics collected. Please qualify this sentence to say metrics are collected up to the limit.
This integration automatically discovers and collects all CloudWatch metrics published for the configured namespaces. The CloudWatch API bills per metric requested, so the more resources that exist in your account, the more metrics are discovered and the higher the collection cost.
packages/aws_cloudwatch_input_otel/docs/README.md:91
- This says the limit defaults to 10,000 for the integration, but the ALB, CLB, NLB, and GWLB policy templates still default to 250 (
manifest.yml:344,404,466,526). Please document the per-stream defaults or update the remaining templates so this cost guidance matches the actual configuration.
To avoid unexpectedly large bills, the number of metrics collected is capped by the **Autodiscover Limit**, which defaults to 10,000, high enough to capture all metrics in a typical deployment. The limit applies to each namespace separately. It is a ceiling, not a fixed value - you are billed only for the metrics that actually exist, so smaller accounts cost proportionally less. With the 10,000 limit, the maximum cost per namespace is ~$3,500 per month at a 5-minute collection interval.
packages/aws_cloudwatch_input_otel/docs/README.md:91
- The stated ~$3,500 maximum does not follow from the documented billing model and the receiver defaults: the manifest requests three statistics per discovered metric, so 10,000 × 3 × 8,640 five-minute intervals × $0.01/1,000 is about $2,592 per 30-day month. Please verify the AWS pricing basis and correct or remove this precise estimate.
To avoid unexpectedly large bills, the number of metrics collected is capped by the **Autodiscover Limit**, which defaults to 10,000, high enough to capture all metrics in a typical deployment. The limit applies to each namespace separately. It is a ceiling, not a fixed value - you are billed only for the metrics that actually exist, so smaller accounts cost proportionally less. With the 10,000 limit, the maximum cost per namespace is ~$3,500 per month at a 5-minute collection interval.
- Files reviewed: 4/4 changed files
- Comments generated: 3
- Review effort level: Lite
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
There was a problem hiding this comment.
🟡 Changes recommended
Documentation and configured defaults are inconsistent, with additional rendering, override, and cost-accuracy issues.
Get a fresh assessment by requesting another Copilot review.
Review details
Suppressed comments (4)
packages/aws_cloudwatch_input_otel/_dev/build/docs/README.md:86
- The character after
###is a non-breaking space (U+00A0), unlike the ordinary ASCII spaces in the surrounding headings. Markdown parsers that require an ASCII space may render this as paragraph text instead of a level-3 heading; replace it in this source and the generated README with a regular space.
### Autodiscover Limit and impact on cost
packages/aws_cloudwatch_input_otel/_dev/build/docs/README.md:88
- This paragraph says the integration collects all metrics, but
discovery.limitexplicitly caps enumeration and the next paragraph says metrics over the cap are dropped. That is misleading for namespaces with more metrics than the configured limit; please say that metrics are collected up to the configured limit.
This integration automatically discovers and collects all CloudWatch metrics published for the configured namespaces. The CloudWatch API bills per metric requested, so the more resources that exist in your account, the more metrics are discovered and the higher the collection cost.
packages/aws_cloudwatch_input_otel/docs/README.md:87
- The separator after
###is a non-breaking space (U+00A0), not an ASCII space. CommonMark/GFM heading syntax may not recognize this as a heading, so the new section can render as a paragraph and be omitted from the table of contents; use a regular space here and regenerate the built README.
### Autodiscover Limit and impact on cost
packages/aws_cloudwatch_input_otel/manifest.yml:105
- These updates leave the Application, Classic, Network, and Gateway ELB policy templates at
default: 250(manifest lines 344, 404, 466, and 526), while the new README says the default is 10,000 and the changelog says the default was increased to 10,000. ELB users therefore retain the old cap and the documented behavior is incorrect; update the remaining service defaults or qualify the documentation and changelog to exclude them.
default: 10000
- Files reviewed: 4/4 changed files
- Comments generated: 3
- Review effort level: Lite
stefans-elastic
left a comment
There was a problem hiding this comment.
Overall the changes look good however Copilot comments looks valid to me and I think they are worth addressing. Apart from inline comments I think it's worth taking a look (and addressing) comments in #21199 (review) (Review details -> Suppressed comments)
|
yes, i agree. i think we could do with adjusting some of the default collection periods here to match the beats version too. e.g. for lamda, sqs and ECS, we collect 1m granularity metrics once every 5m; this is still only billed as one metrics request, so it keeps the cost aligned with the documented example. i'll make the changes. this should all improve dramatically once open-telemetry/opentelemetry-collector-contrib#50274 is merged too, since we won't need to collect all the stats for each metric, so we should see a 3-4x reduction in the number of metrics requested across the board. |
… cost ceiling; clarify README.
There was a problem hiding this comment.
🟡 Changes recommended
The documented autodiscover_limit setting is hidden from the UI, and missed-window behavior needs clearer qualification.
Get a fresh assessment by requesting another Copilot review.
Review details
Suppressed comments (3)
packages/aws_cloudwatch_input_otel/_dev/build/docs/README.md:74
- This wording promises no data loss, but the receiver derives each query window from the current time and does not persist a cursor or backfill a missed window. A failed scrape or collector restart can therefore leave a gap (and the new 5-minute interval makes that gap larger); please qualify this as applying to uninterrupted collection and mention the missed-window behavior.
- Set **Collection Interval** to **Period**, or to a multiple of it. Each poll returns every data point in the elapsed window, so polling less often reduces cost proportionally without losing data — it only delays how soon the data arrives. Polling more often than Period re-reads the same data point and multiplies cost for no benefit.
packages/aws_cloudwatch_input_otel/_dev/build/docs/README.md:95
autodiscover_limitis markedshow_user: falsein every policy template (for example, manifest.yml:101-107), so it is not exposed in the integration's advanced settings. Users cannot follow this instruction through the UI; either expose the variable or document a supported policy/API override instead.
If you need to reduce cost, you can lower the **Autodiscover Limit** in the integration's advanced settings. Reducing the limit below the number of metrics in your account means some metrics will not be collected, and which ones are dropped is not predictable. We recommend changing this value only if you understand the metric volume in your AWS account, and lowering it gradually while verifying that the metrics you rely on are still present.
packages/aws_cloudwatch_input_otel/docs/README.md:75
- This wording promises no data loss, but the receiver derives each query window from the current time and does not persist a cursor or backfill a missed window. A failed scrape or collector restart can therefore leave a gap (and the new 5-minute interval makes that gap larger); please qualify this as applying to uninterrupted collection and mention the missed-window behavior.
- Set **Collection Interval** to **Period**, or to a multiple of it. Each poll returns every data point in the elapsed window, so polling less often reduces cost proportionally without losing data — it only delays how soon the data arrives. Polling more often than Period re-reads the same data point and multiplies cost for no benefit.
- Files reviewed: 4/4 changed files
- Comments generated: 1
- Review effort level: Lite
stefans-elastic
left a comment
There was a problem hiding this comment.
I've done a little test with RDS policy template:
The autodiscover limit is correct in the rendered policy:

data is getting collected successfully and each 1-minute datapoint has consistent amount of records (which I believe means there is no data dropped):

5 last minutes are always empty (it's because of Delay, right?)
I wasn't able to test the limit properly (as I am unsure where can I get that much data)
|
yeh exactly. |
There was a problem hiding this comment.
🔵 Needs a closer look
Documentation service listings and cost assumptions need correction before approval.
Review details
Suppressed comments (3)
packages/aws_cloudwatch_input_otel/_dev/build/docs/README.md:93
- The ~$2,600 figure is only the three-stat case, not an unconditional maximum: the manifest defaults to two statistics for RDS/SQS and three for the other templates, while
autodiscover_aggregationcan request additional statistics and the interval can be overridden. Please state that the estimate assumes three statistics at the default 5m interval, and that additional statistics or a shorter interval increase the cost.
To avoid unexpectedly large bills, the number of metrics collected is capped by the **Autodiscover Limit**, which defaults to 10,000, high enough to capture all metrics in a typical deployment. The limit applies to each namespace separately. It is a ceiling, not a fixed value - you are billed only for the metrics that actually exist, so smaller accounts cost proportionally less. With the 10,000 limit, the maximum cost per namespace is ~$2,600 per month at the default 5-minute collection interval.
packages/aws_cloudwatch_input_otel/_dev/build/docs/README.md:88
- This generated copy also lists Classic, Network, and Gateway ELB in the defaults table while the earlier “Supported services” table omits them. Once the source README is corrected, regenerate this file so the complete service inventory is consistent here too.
| AWS Classic ELB | 5m | 1m |
| AWS Network ELB | 5m | 1m |
| AWS Gateway ELB | 5m | 1m |
| AWS ECS | 5m | 1m |
packages/aws_cloudwatch_input_otel/docs/README.md:87
- The new defaults table now lists Classic, Network, and Gateway ELB, but the earlier “Supported services” table still lists only Application ELB and ECS/Fargate. Since that section presents the complete service list, users can miss these supported templates; add the three ELB variants there as well.
| AWS SQS | 5m | 1m |
| AWS Application ELB | 5m | 1m |
| AWS Classic ELB | 5m | 1m |
| AWS Network ELB | 5m | 1m |
| AWS Gateway ELB | 5m | 1m |
- Files reviewed: 4/4 changed files
- Comments generated: 0 new
- Review effort level: Lite
There was a problem hiding this comment.
🟡 Changes recommended
Documentation issues remain, including one moderate cost-estimate clarification.
Get a fresh assessment by requesting another Copilot review.
Review details
Suppressed comments (2)
packages/aws_cloudwatch_input_otel/_dev/build/docs/README.md:93
- The ~$2,600 figure is described as a maximum per namespace, but each policy is configured for a specific AWS region and the GetMetricData cost is incurred per region/namespace. Deployments that collect the same namespace in multiple regions (or enable multiple service policies) can therefore exceed this aggregate figure; please label it as a per-region estimate and explain that totals sum across policies.
To avoid unexpectedly large bills, the number of metrics collected is capped by the **Autodiscover Limit**, which defaults to 10,000, high enough to capture all metrics in a typical deployment. The limit applies to each namespace separately. It is a ceiling, not a fixed value - you are billed only for the metrics that actually exist, so smaller accounts cost proportionally less. With the 10,000 limit, the maximum cost per namespace is ~$2,600 per month at the default 5-minute collection interval.
packages/aws_cloudwatch_input_otel/manifest.yml:466
- Now that the default is 10,000, the NLB-specific text is stale: 200+ combinations do not by themselves justify increasing the limit above 10,000. Please update this description to explain when a value above the new default is actually needed.
description: >-
Maximum number of metrics to enumerate for this namespace.
NLB can have 200+ metric/dimension combinations — set this higher than the default if metrics are missing.
default: 10000
- Files reviewed: 4/4 changed files
- Comments generated: 1
- Review effort level: Lite
|
✅ All changelog entries have the correct PR link. |
There was a problem hiding this comment.
🟢 Approval recommended
Only a non-blocking documentation nit remains.
Review details
Suppressed comments (1)
packages/aws_cloudwatch_input_otel/_dev/build/docs/README.md:77
- The unchanged
delay: 5mmeans this new 5m interval queries roughly the window from now-10m to now-5m for the 1m-period services. Although no points are lost, the worst-case age of data in Elasticsearch increases to about 10 minutes; please document this latency trade-off alongside the cost savings so users do not interpret the change as only a polling/cost detail.
- Set **Collection Interval** to **Period**, or to a multiple of it. Each poll returns every data point in the elapsed window, so polling less often reduces cost proportionally without losing data — it only delays how soon the data arrives. Polling more often than Period re-reads the same data point and multiplies cost for no benefit.
- Files reviewed: 4/4 changed files
- Comments generated: 0 new
- Review effort level: Lite
💚 Build Succeeded
History
|
|
Package aws_cloudwatch_input_otel - 0.7.0 containing this change is available at https://epr.elastic.co/package/aws_cloudwatch_input_otel/0.7.0/ |
Proposed commit message
increase cloudwatch autodiscover limit and add section to readme explaining the consequence
Checklist
changelog.ymlfile.Author's Checklist
How to test this PR locally
Related issues
Screenshots